Lightweight, highly optimized CPU runtime for GGUF models and embeddings.
Private, local AI desktop app — run open LLMs fully offline with a coding agent, knowledge base and voice.